Back

Applications in Plant Sciences

Wiley

Preprints posted in the last 30 days, ranked by how well they match Applications in Plant Sciences's content profile, based on 23 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Making the most out of it: shallow genome-skimming possibilities for the systematics of prickly lineages of Solanum (Solanaceae)

Alves, R. T. d. L.; Gouvea, Y. F.; Dalapicolla, J.; Poczai, P.; Giacomin, L. L.

2026-07-09 plant biology 10.64898/2026.07.08.737304 medRxiv
Top 0.1%
37.9%
Show abstract

Premise: Genome skimming (GS) is a cost-effective approach for plant phylogenomics, but its ability to recover informative datasets from different genomic compartments, particularly genome-wide SNPs, remains poorly explored in Solanum. Methods: We evaluated shallow GS for phylogenetic inference in South American prickly Solanum lineages by recovering plastid, mitochondrial, and nuclear datasets, including coding regions and genome-wide SNPs. Phylogenies were inferred using maximum-likelihood and coalescent approaches under different SNP filtering strategies. Results: GS successfully recovered complete plastomes, organellar coding regions, and large SNP datasets, but failed to consistently assemble mitochondrial genomes or recover low-copy nuclear genes. SNP-based analyses, especially from the nuclear genome, produced stable, well-supported phylogenies that were largely congruent across inference methods. In contrast, coding-region datasets, particularly from the mitochondrial genome, showed greater topological discordance, revealing cytonuclear conflict. Discussion: Our results demonstrate that shallow GS is an effective strategy for generating informative SNP datasets for phylogenetic inference in Solanum, despite limitations in recovering complete mitochondrial genomes and low-copy nuclear loci. SNP-based analyses substantially expand the phylogenetic potential of GS, providing a practical and cost-effective alternative for systematic studies.

2
SeedMeasure: an efficient approach and open-source program to quantify seed size

Sims, B.;Gaudinier, A.;Blackman, B.

2026-06-29 Plant Biology 10.64898/2026.06.27.734974 medRxiv
Top 0.1%
10.1%
Show abstract

PremiseSeed size and morphology are critical traits in agriculture, ecology, and genetics, but high-throughput quantification of these traits is often limited by labor-intensive manual measurements or expensive, platform-specific imaging software. Methods and ResultsWe developed SeedMeasure, a lightweight, open-source, and cross-platform command-line tool written in Python that automates the measurement of seed area, length, and width from images. Using a simple imaging setup, the program processes images by correcting for perspective skew, filtering debris, and exports quantitative data alongside quality-check images. We validated SeedMeasure across nine diverse species, ranging from small Arabidopsis thaliana seeds to large Zea mays kernels. The tool quickly handles images using multithreading and demonstrates high reproducibility, yielding low coefficients of variation across repeated runs. ConclusionsCompared to existing software, SeedMeasure is free, offers faster processing through parallel computing, and provides standalone executables that require no programming dependencies. SeedMeasure offers an accessible, cost-effective, and high-throughput approach for rapid phenotypic profiling, making advanced seed morphological analysis available to researchers without specialized laboratory hardware.

3
EcoMorph: Universal morphological trait quantification from natural language prompts for ecological research

Amoah, E. I.; Bunch, Z.; Thomas, H. M.; Patch, H. M.; Grozinger, C.

2026-07-12 bioinformatics 10.64898/2026.07.10.737871 medRxiv
Top 0.1%
7.2%
Show abstract

0.O_LIMorphological traits such as floral area and body size are fundamental to ecological research, serving as inputs for studies of pollinator-plant interactions, habitat quality, and biodiversity monitoring. However, accurately measuring these traits from images remains challenging, particularly in complex field conditions where existing tools exhibit reduced accuracy and limited generalizability across taxa. C_LIO_LIWe present EcoMorph, a modular morphological measurement system that leverages the Segment Anything Model 3 (SAM3) to quantify traits across diverse ecological contexts. Unlike task-specific segmentation models requiring domain-specific training data, SAM3s prompt-based architecture enables segmentation of arbitrary biological structures from natural-language prompts, using the same underlying model across flowers, insects, and other targets without retraining. From the resulting segmentations, EcoMorph extracts three classes of measurement: area, linear dimensions, and object counts. C_LIO_LIWe validated EcoMorph across two ecological scales. At the intermediate scale, EcoMorph-derived floral area agreed closely with manual ImageJ measurements (R2 = 0.935, n = 74) under simple-background conditions and (R2 = 0.928, n = 58) under complex-background conditions, with valid predictions for 95% of images. At the fine scale, EcoMorph-derived insect body area was strongly correlated with hand-measured intertegular distance (r = 0.810, n = 349), capturing body-size variation across species from the small Bombus impatiens to the large Xylocopa virginica. Object counts matched manual counts almost exactly for well-separated insects in an insect box (R2 = 0.9997, n = 12). C_LIO_LIBy combining prompt-based segmentation with modular measurement, EcoMorph enables high-throughput quantification of area, size, and abundance from heterogeneous image sources without taxon-specific training. This generality supports a broad range of ecological applications, including pollinator and plant trait research, biodiversity and abundance monitoring, and allometric biomass estimation. C_LI

4
A Practical Roadmap For Sampling Floral Nectar From Communities of Many Plant Species

Kirschke, G. E.; Bain, J. A.; Ogilvie, J. E.; CaraDonna, P. J.

2026-06-23 ecology 10.64898/2025.12.19.695174 medRxiv
Top 0.1%
6.7%
Show abstract

O_LIFloral nectar plays a critical role in shaping the ecology and evolution of plant-pollinator interactions. Effective and efficient methods that allow for broad-scale sampling of nectar volume and sugar concentration across a diversity of taxa are needed to improve our understanding of many dimensions of mutualistic plant-pollinator interactions--including their basic ecology and evolution, their responses to environmental change, and their conservation and restoration. C_LIO_LIDespite the key importance of nectar for mediating plant-pollinator interactions, quantifying floral nectar in the field from many different plant species is challenging because there is often no one-size-fits-all sampling method that is effective across a diversity of floral structures and nectar traits. Different methods require different preparation, and sampling from many species involves a variety of logistical challenges. C_LIO_LIHere we provide a methodological roadmap for sampling floral nectar in the field from many different plant species. We describe our nectar collection methods in detail, including necessary equipment, calculations, and approaches appropriate for different floral morphologies. We also provide a troubleshooting guide for common problems encountered while collecting nectar in the field. To demonstrate the utility and effectiveness of our methods for collecting nectar from many different species, we present results on nectar trait variation from 53 species in an ecosystem. C_LIO_LIOur method illustrates that nectar traits vary considerably within and among plant species, indicating that large-scale nectar sampling projects are an important consideration for many basic and applied questions in pollination ecology and evolution. We hope that across many plant communities and ecosystems, our paper provides a practical roadmap for how to navigate the complexities of quantifying floral nectar traits. C_LI

5
UVfinder: a tool to extract bryophyte sex-linked gene copies from the GoFlag408 probe set

Kim, S.; Bowman, J.; Braun, E. L.; McDaniel, S.

2026-07-07 bioinformatics 10.64898/2026.07.01.735932 medRxiv
Top 0.1%
6.6%
Show abstract

Target enrichment sequencing using probe sets like GoFlag 408 has revolutionized phylogenetics, yet recent genomic data indicate that some probes may be sex-linked, potentially introducing topological conflict while also allowing studies of sex-specific evolutionary processes. To test for sex-linkage across the bryophytes, we developed UVfinder, a pipeline designed to identify sex-linked GoFlag loci across published moss genomes and enable sex-aware downstream analyses. Applying UVfinder to 50 dioicous moss genomes, we identified 93 probes that exhibit sex-linkage in one or more lineages, providing genomic evidence for neo-sex chromosome formation via autosome-sex chromosome fusion and gene translocation. Furthermore, by comparing species trees derived from sex-linked versus autosomal loci in Hypnales and Dicranidae, we demonstrate that sex-linked loci harbor phylogenetic information that is distinct from that in autosomes. We also discovered a pervasive female sampling bias in the genomic data, perhaps reflecting a preference among collectors for plants with sporophytes. Ultimately, our findings highlight the dynamism in sex linkage across bryophytes and suggest that sex-aware phylogenomics can be used to reconstruct ancestral karyotypes and potentially resolve topological conflict. We expect that UVfinder will facilitate the further study of sex-specific evolutionary processes, particularly with improved genome assemblies and increased sampling in males.

6
Small representative samples can capture global vascular plant diversity patterns

Baldaszti, L.; Moonlight, P.; Brummitt, N.; Pironon, S.; Sarkinen, T.

2026-07-10 plant biology 10.64898/2026.07.08.737287 medRxiv
Top 0.1%
6.2%
Show abstract

Incomplete information on distributions for a high proportion of the world's plant species together with biases in global biodiversity data mean that current estimates of plant diversity patterns are skewed. A key issue is that current predictions rely on a subset of species that is not representative of all plant species. Here we tested the feasibility of a representative sampling approach for mapping global vascular plant diversity at the finest scale where comprehensive data is available. Using the World Checklist of Vascular Plants as a reference, we generate random samples of species with increasing sample sizes from the global species pool. We compare the diversity patterns retrieved from the samples against the patterns of the reference dataset using spatially weighted correlation coefficients and four different diversity metrics. We find that at the botanical country scale, representative global maps of species and phylogenetic diversity can be created with small numbers of species (~1% [0.2% and 0.4%, respectively]) at the botanical country scale. For effective growth form and family diversity sample sizes encompassing ~20% [19.2% and 19.5%, respectively] of all species are needed. Random samples require markedly fewer species to reach high correlations than when restricting the pool of species to single plant families or genera. We show that when representative samples are used robust inferences of plant diversity patterns can be made from only a small proportion of species.

7
Diversity Assessment with SNP, SSR, AFLP, and RAPD Markers in Plants: A Systematic Review and Meta-Analysis

Olagunju, Y. O.; Olawuyi, O. J.

2026-07-07 plant biology 10.64898/2026.07.03.736291 medRxiv
Top 0.1%
6.2%
Show abstract

Background. DNA-based molecular markers underpin plant genetic diversity assessment, germplasm characterisation, and conservation prioritisation. Four marker systems dominate the field: Amplified Fragment Length polymorphisms (AFLPs), simple sequence repeats (SSRs), single nucleotide polymorphisms (SNPs), and random amplified polymorphic DNA (RAPDs). No quantitative meta-analysis had pooled their performance on the canonical diversity metrics: polymorphism information content (PIC), expected heterozygosity (He), and resolution power, across plants. Existing reviews are narrative, marker-restricted, or qualitatively conclusive of infeasibility. Methods. A PRISMA 2020-compliant systematic review (registered at the Open Science Framework) was executed. Eligible studies were within-study paired comparisons genotyping the same accession panel with at least two of {SNP, SSR, AFLP, RAPD} and reporting at least one diversity metric. Effect sizes were paired standardised mean differences (Hedges' g) computed under the Bernoulli-variance approximation. Random-effects REML meta-analysis used metafor 5.0.1 with Knapp-Hartung adjustment, leave-one-out, and r-sensitivity. Results. Fifteen within-study paired contrasts were eligible, distributed across three pools. Pool 2 (SSR vs SNP, He, k = 5) yielded a pooled Hedges' g of 0.494 (95% CI: -0.078 to 1.066, p = 0.075; I-squared = 90.2%; 95% PI [-0.82, 1.81]). SSRs exceeded SNPs on He in 4 of 5 studies; leave-one-out removal of the panel-size-asymmetric outlier raised the estimate to g = 0.644 (p = 0.025). Pool 3a (dominant-marker stratum, k = 6) yielded g = 0.419 (95% CI: -0.121 to 0.960, p = 0.103; I-squared = 56.5%); five of six contrasts showed SSR or AFLP exceeding RAPD on per-locus PIC. Pool 1 (PIC, k = 3, exploratory) gave a consistent direction (g = 0.453). All three pools point in the same direction: codominant or AFLP markers carry more per-locus information than the alternative being compared. Conclusions. SSR markers reported higher per-locus diversity than SNP and RAPD markers in plant within-study paired comparisons, mechanistically grounded in the SNP biallelic ceiling and the multi-allelic richness of SSRs. The effect attenuated or reversed in selfing/low-diversity panels and at the per-panel level when SNP panels exceeded approximately 1000 loci. RAPDs show the lowest per-locus information content of the four classes.

8
Near-Gapless and Haplotype-Resolved Capsella Genomes Enable Investigation into Genomic Consequences of Mating System Shifts

Chen, H.; Emmerson, R.; Mosher, R. A.

2026-07-10 plant biology 10.64898/2026.07.10.737683 medRxiv
Top 0.1%
5.3%
Show abstract

The shift from outcrossing to self-fertilization is a common evolutionary transition in flowering plants. The genus Capsella, comprising the obligate outcrosser C. grandiflora and two self-fertile species, C. rubella and C. orientalis, provides a powerful system to explore genomic consequences of mating system shifts. Despite its utility, existing genomic resources in Capsella are fragmented, incomplete, and particularly deficient in repetitive genomic regions, hindering the study of transposable element (TE) dynamics and gene annotation. Here, we present high-quality, chromosome-scale, near-gapless genome assemblies for C. grandiflora, C. rubella, and C. orientalis. Leveraging these improved genomes, we created high-quality genomic resources for the Capsella genus by performing comprehensive, de novo annotations of protein-coding genes and TEs. Comparative genomic analysis among these species reveals differences in TE abundance, position, and production of small RNAs. These resources provide an unprecedented opportunity to explore how mating system transitions influence genome architecture, TE behavior, and gene evolution. This research also developed a static online platform for Capsella genomic resources, Capsella Database (CapBase, www.capsella.uk), to support community use of these resources. Our findings advance understanding of the genomic impacts of selfing and establish a robust foundation for future research into genomics, epigenomics, and evolutionary biology within Capsella and related plant systems.

9
An evaluation of clustering and assembly strategies from Iso-Seq data in the absence of reference genomes in non-model animals

Eleftheriadi, K.; Vazquez-Valls, M.; Fernandez, R.

2026-07-08 evolutionary biology 10.1101/2025.09.18.677004 medRxiv
Top 0.1%
4.7%
Show abstract

Transcriptome assembly enables the recovery of expressed genes and isoforms, but the optimal strategy for reconstructing transcriptomes from long-read sequencing remains unresolved. In particular, establishing best practices for generating accurate gene models and selecting representative isoforms is essential for comparative genomics, as for orthology inference typically only the longest isoform per gene model is included. Here, we systematically compare clustering and de novo assembly methods using PacBio Iso-Seq data from 17 animal lineages spanning seven phyla, most of them non-model species, with the goal of investigating which methodology is more adequate to select one isoform per gene model, in the absence of specific pipelines to do so. We evaluate four approaches: isoseq cluster, CD-HIT, RNA-Bloom2 and isONform. We benchmark them with short-reads using Trinity, assessing assembly quality with BUSCO completeness, short-read mapping rates, coding sequence recovery, and longest isoform prediction. Our results show that CD-HIT clustering at high similarity thresholds ([≥]99%) yields the most complete and coding-rich long-read transcriptomes, rivaling Trinity while avoiding its high redundancy. Consensus-based methods such as isoseq cluster and isONform recover fewer single-copy orthologs (mirrored in a lower BUSCO score) and achieve lower mapping rates, while RNA-Bloom2 provide intermediate performance with reduced duplication. Together, these findings establish, to date, CD-HIT as a robust and practical strategy for transcriptome reconstruction from long-read data when genomic references are unavailable. By benchmarking de novo methods across a taxonomically broad dataset, this work defines the realistic capabilities of long-read transcriptome reconstruction in the absence of a reference genome and provides practical guidance for deriving high-quality gene models and selecting representative isoforms for orthology inference in non-model species.

10
trAIt: Species-by-Trait Data Retrieval using Large Language Models

Balaji, S.; Martinson, K. A.; Schellenberger, J. S.; Koley, J.; Inman, C. M.; Hofmann, H. A.; Young, R. L.; Harpak, A.

2026-06-24 bioinformatics 10.64898/2026.06.19.732660 medRxiv
Top 0.1%
3.3%
Show abstract

Biological research often requires information about species traits. Manual literature collation can be time-consuming and miss parts of the literature. To address this gap, we developed trAIt, a publicly available software for the retrieval of characteristics of species from scientific literature catalogued in the Europe PubMed Central (PubMed) database. trAIt provides a graphical user interface (GUI) in which users specify species and characteristics of interest. Leveraging a large language model (LLM), trAIt retrieves relevant papers, combines their content through a consensus-based summarization model, and outputs a species-by-characteristic table. For a case study involving frog species, trAIt recovered 47.1% of trait-species combinations in 2.75 hours, while an expert curator independently recovered 62.4% over months. The consensus-based summarization substantially aids accuracy compared to single-source extraction. Across three case studies of vertebrate taxa, an expert confirmed the accuracy of 70.9% of trait-species entries recovered by trAIt. We observed considerable variation across taxa in trAIts accuracy, which is possibly due to heterogeneity in open-access literature availability and inconsistencies in species and trait terminology. In sum, our analysis suggests that LLM-based tools can accelerate biological data synthesis but should be used to support domain experts research, rather than replace their judgment.

11
Large phenological advances and delays over 124 years of climate change alter co-flowering among North American Viola

Edwards, C.; Moyle, L. C.

2026-07-02 evolutionary biology 10.64898/2026.06.27.734973 medRxiv
Top 0.1%
2.7%
Show abstract

Shifts in flowering phenology are one of the most well studied plant responses to global climate change. Many studies have documented these shifts and their drivers, including some that describe altered patterns of co-flowering among taxonomically broad species within communities. In comparison, few analyses have examined systematic changes in co-flowering between closely related, interfertile species, where co-flowering can have unique evolutionary consequences. To address such shifts in co-flowering among close relatives, we investigate phenological responses to climate change and its effect on patterns of co-flowering over the past 124 years in 52 species of North America violets (Viola). This genus has many co-occurring species that reproductively interact via shared pollinators and hybridization. We use ~14,000 herbarium records along with environmental and species trait data to model the magnitude of recent flowering phenology shifts, environmental variables and/or species traits associated with these shifts, and resulting changes in co-flowering among species. While both the magnitude and direction of phenological shifts varied among Viola species, nearly half (25/52) show significant changes in flowering day. Regardless of whether flowering was advanced or delayed, flowering date was most consistently associated with local mean temperature. Of six species-level traits, geographical region also significantly predicted flowering shifts, consistent with environment and geography together explaining broad phenological responses across this group. These shifts have produced significant changes to pairwise patterns of co-flowering among species -- ranging from a 59 day increase in co-flowering to complete loss of co-flowering overlap. Sympatric pairs specifically have experienced both increases and decreases in co-flowering, with a geographic pattern of increased co-flowering occurring mainly in eastern US and decreased co-flowering common in western US. Because Viola species are generalist pollinated and already known to hybridize, these new co-flowering patterns could further undermine reproductive barriers among species in this genus.

12
Griphus Software for Multi Panel Figure Composition and Experimentation with Emphasis on Taxonomy

Aguiar, A. P.

2026-07-11 zoology 10.64898/2026.07.07.736512 medRxiv
Top 0.1%
2.4%
Show abstract

The preparation of multi panel figures remains a labor intensive step in scientific publication. Albeit there are specific tools available to solve this problem, they are often highly specialized, difficult to install, or time consuming to learn. Griphus is a standalone graphical application designed for rapid composition and experimentation with multi panel figures, developed by and for zoological taxonomists. Functions specifically designed for multi panel composition include automatic figure numbering and placement, aspect ratio operations, spacers, layout rotation, layout suggestions, and automatic generation of figure legends, including scale bar descriptions. The software can perform both spatial interpretation of images on the canvas and work with a simple, editable layout formula. It also enables instant multi panel composition, with numbered images and automatic contrast selection for the numbers, obtained simply by loading images. User defined parameters such as target printable dimensions, resolution, spacing, and color mode are preserved throughout the work. The program produces coordinated outputs consisting of the final composite figure, a readable file describing the layout structure, and a .gri file storing images, transformations, and parameters for exact regeneration. Griphus is intended as a complementary tool to professional image software, providing a simple and efficient environment for constructing high quality multi panel figures.

13
A Highly Contiguous Reference Genome for Scalesia gordilloi (Asteraceae), a Critically Endangered Plant Endemic to the Galapagos Islands

Pozo, G.; Rivas-Torres, G.; Velez-Darquea, E.; Barragan-Orbe, D.; Torres, M. d. L.

2026-06-29 genomics 10.64898/2026.06.25.734018 medRxiv
Top 0.2%
2.3%
Show abstract

Scalesia gordilloi is a critically endangered species endemic to San Cristobal Island in the Galapagos archipelago and represents one of the most unique and vulnerable lineages within the adaptive radiation of the genus Scalesia. Despite its evolutionary distinctiveness and conservation importance, no genomic resources have been available for this species. Here, we present the first high-quality reference genome of S. gordilloi, generated using Oxford Nanopore long-read sequencing. Across three PromethION R10.4.1 flow cells, we obtained 80.5 Gb of long reads (~25X coverage), which enabled a highly contiguous 3.61 Gb assembly composed of only 549 contigs and an N50 of 106.6 Mb. BUSCO completeness reached 98.6%, with assembly metrics comparable to other high-quality Asteraceae genomes. Repeat annotation revealed that 76.2% of the genome is composed of interspersed elements, dominated by LTR retrotransposons. Structural annotation resulted in 47,913 high-confidence protein-coding genes, consistent with expectations for large, repetitive Asteraceae genomes. This genome provides a critical foundation for conservation genomics, enabling assessments of genetic diversity, inbreeding, and adaptive potential in the species. It further establishes a framework for comparative genomics across the Scalesia radiation and supports future efforts to protect and restore one of the most threatened plant lineages of the Galapagos Islands.

14
DeepPheno: A Deep Learning Framework for Linking Hyperspectral Imaging and SNP Genotypes in Lettuce

Okyere, F. G. G.; Mehrem, S. L.; Snoek, B. L.; Van den Ackerveken, G.; Abeln, S.

2026-07-10 plant biology 10.64898/2026.07.09.737449 medRxiv
Top 0.2%
2.1%
Show abstract

While whole genome sequencing captures millions of single nucleotide polymorphisms (SNPs) and hyperspectral imaging (HSI) enables non destructive plant phenotyping, integrating these modalities to link genotype to phenotype remains challenging due to their high dimensionality and non linearity. This study presents DeepPheno a deep learning framework that predicts SNP genotypes from HSI data, using model predictability as a proxy for genotype phenotype association. HSI data were acquired from 194 lettuce genotypes under field conditions. HSI data patches (20 x 20 pixels x 224 spectral bands) were used to train a hybrid CNN to predict the variant of a specific SNP. The framework was validated on SNPs with known phenotypic effects (anthocyanin, leaf serration, pale pigmentation), achieving high predictive performance (AUC ranging from 0.806 to 0.935), whereas models trained on randomly shuffled labels performed at chance (mean AUC {approx} 0.51). Extending the workflow to 50 randomly selected putatively neutral SNPs, most yielded low predictability, but two showed high performance (AUC > 0.76), suggesting uncharacterized genotype phenotype links. Explainable AI, including SHAP and Grad CAM, identified relevant spectral and spatial features driving these predictions, particularly the green and red edge wavelengths associated with pigment dynamics and leaf structure. These results establish a framework for understanding complex genotype phenotype interactions in plants and extracting these links from HSI data without predefining the exact trait values. It provides an avenue for high throughput trait discovery and description and extends the integration of image based phenomics with plant genetics.

15
From mountaintops to metacollections: using genomics to evaluate ex situ conservation collections. A case study from tropical montane cloud forest plants

Cascini, M.; Simpson, L.; Worboys, S.; Worboys, W.; Guja, L.; Knapp, Z.; Bredell, P.; Percival, J.; Rossetto, M.; Crayn, D.

2026-07-03 genomics 10.64898/2026.06.27.734930 medRxiv
Top 0.2%
1.5%
Show abstract

A core aim of ex situ conservation is to represent wild genetic diversity in managed living collections. For the climate-threatened tropical montane cloud forest (TMCF) flora of northeast Australia, an ex situ metacollection of plants and seeds has been established by the Tropical Mountain Plant Science (TroMPS) project. In this study we used reduced-representation sequencing (DArTseq) of wild, herbarium, and ex situ material alongside provenance information for ten species, to pursue two central aims: to characterise landscape-scale genetic structure across species' ranges, and to evaluate how well the assembled metacollections represent that wild diversity. Analyses revealed consistent patterns of genetic differentiation among mountain top populations across multiple species, reflecting the isolating influence of lowland gaps between upland habitats, with the degree of differentiation varying among species. These results provide the first genetic baseline for Australian TMCF flora and reinforce the importance of treating individual mountain top populations as distinct units for conservation management. Additionally, the project provided valuable insights into the logistical challenges of coordinated multi-institutional collecting, informing strategies for metacollection design more broadly. Evaluation of the metacollection revealed both strengths and gaps in representation across species, providing an evidence base to refine the current holdings and guide future targeted collecting to strengthen their long-term conservation value.

16
Seasonal climatic impacts on orchid productivity in an urban ecosystem

Brundrett, M.

2026-06-30 Plant Biology 10.64898/2026.06.29.735162 medRxiv
Top 0.2%
1.5%
Show abstract

ContextThe global diversity hotspot in Southwest Australia has >480 orchids facing increasing threats from climate extremes, fire and habitat decline. AimsTo develop effective and consistent tools for measuring climate impacts on productivity in a diverse urban orchid community. MethodsAnnual variations in flower and seed production for 17 orchids were determined using thousands of records over a decade with extreme climate variability. Key resultsRainfall deficits and temperatures in autumn, winter and spring increased substantially over 125 years. Seasonal climate anomalies reduced flowering and seed production for orchids, but this varied between species and seasons. These effects were summarised by climate response (CRI) and sensitivity (CSI) indexes. Early or late flowering species were most vulnerable to seasonal drought, and visually deceptive pollination preferred warm dry conditions. CRIs were strongly correlated with orchid pollination syndromes and flowering times. Effects on mycorrhizal fungi and pollinators were also observed. Extrapolating climate trends to 2100 predicted further impacts on orchid productivity (-5-40%). ConclusionsOrchid climate responses were diverse and deeply integrated with pollination, phenology, fire sensitivity and other key traits. ImplicationsResearch in an urban climate observatory produced a climate analysis framework that is likely relevant to many orchids and other biota.

17
simSOMA: a cell-lineage based simulator of the somatic VAF spectrum in plants

Johannes, F.

2026-07-01 genomics 10.64898/2026.06.28.735079 medRxiv
Top 0.3%
1.3%
Show abstract

Plants accumulate somatic mutations during growth, and some of these mutations can spread from local cell lineages into branches, organs, or reproductive tissues. There is growing interest in these variants because they can underlie bud-sport traits in crops, contribute to within-organism somatic selection, and provide genetic variation that may be transmitted vegetatively or sexually to future generations. Recent genomic sequencing of bulk and layer-enriched plant tissues has shown that de novo somatic variants can generate complex variant allele-frequency (VAF) spectra. Interpreting these spectra requires understanding how mutations arising during mitotic cell division are filtered or amplified through shoot growth, branching, and organ formation. Because these processes interact across multiple scales, their combined effects are difficult to derive analytically. Here, we present simSOMA, a modular simulator that links rooted plant topologies to explicit cell-lineage dynamics. simSOMA models somatic mutation accumulation during stem-cell self-renewal in the shoot apical meristem, clonal expansion from the stem-cell niche to the meristem periphery, branch founding, and organ formation. Applying simSOMA across diverse growth scenarios revealed how individual processes can be isolated, varied, and combined to assess their effects on organ-level VAF spectra and among-organ variant sharing. The same simulated spectra can also be transformed to represent bulk or layer-enriched sampling and phased or unphased variant readouts, separating effects of developmental history from those introduced by tissue composition and allele counting. Because simSOMA is organized around modules with defined input-output interfaces, individual developmental components can be replaced or extended as new empirical information becomes available. This makes simSOMA a flexible tool for testing alternative models of somatic mosaicism in plants and for guiding the design and interpretation of VAF-based sequencing studies. The simulator is available at https://github.com/jlab-code/simSOMA.

18
Knowledge-guided Bayesian optimization using pre-trained LLMs speeds up the identification of superior genotypes from germplasm collection

Hamazaki, K.; Tsuda, K.

2026-07-02 bioinformatics 10.64898/2026.06.28.735149 medRxiv
Top 0.3%
1.1%
Show abstract

Background: Germplasm collections contain wide genetic diversity that is valuable for plant breeding, but conducting phenotypic evaluation for all genotypes in field trials is rarely feasible. Bayesian optimization offers a way to decide, season by season, which genotypes to cultivate in order to identify superior genotypes with fewer evaluations. However, standard Bayesian optimization commonly starts from randomly selected genotypes and mainly relies on surrogate models built from marker genotype information, while the text-based passport information that accompanies germplasm is not fully used. We examined whether pre-trained large language models can provide prior knowledge that improves these decisions in germplasm evaluation. Results: We constructed a large-language-model-guided Bayesian optimization framework that introduces large language models into two parts of the Bayesian optimization workflow. In zero-shot warmstarting, a large language model proposes initial genotypes using passport information such as cultivar name, country of origin, and subpopulation, optionally together with principal component scores derived from genome-wide single-nucleotide-polymorphism markers. In addition, we evaluated a large-language-model-based surrogate model that predicts phenotypic values for untested genotypes using in-context learning from previously evaluated genotypes. Using a rice germplasm panel and two target traits (seed number per panicle for maximization and protein content for minimization), we compared strategies. For seed number per panicle, zero-shot warmstarting with a general-purpose instruction-following model reduced the number of evaluated genotypes needed to reach the best genotype, whereas improvements were small for protein content. When genomic information was available, Gaussian-process-based Bayesian optimization was the strongest overall approach, while the large-language-model-based surrogate model outperformed random baselines and was competitive in some settings. When genomic information was not available, predictions based on passport information improved efficiency compared with fully random strategies. Conclusions: Pre-trained large language models can inject useful agronomic knowledge into Bayesian optimization for germplasm evaluation, particularly by improving early-stage genotype selection, and can also support optimization when genomic information is unavailable. As models better handle long genomic sequences together with passport information, large-language-model-guided Bayesian optimization may become a practical and explainable decision-support approach for agricultural optimization.

19
GBZ-base and GAF-base: Indexed pangenome file formats

Siren, J.; Paten, B.; the Human Pangenome Reference Consortium,

2026-07-11 bioinformatics 10.64898/2026.07.10.737775 medRxiv
Top 0.3%
1.1%
Show abstract

MotivationExisting pangenome file formats are designed for batch processing. Graphs must be loaded into memory, and alignment files must be read sequentially. Indexed file formats that can be used directly from disk would be more appropriate for interactive applications. ResultsWe propose GBZ-base and GAF-base -- SQLite-backed file formats comparable to GBZ and GAF. GBZ-base supports efficient extraction of local subgraphs, and GAF-base lets us extract all alignments to the subgraph. Additionally, GAF-base is smaller than any other file format for sequence-to-graph alignments. Availability and implementationFrom https://github.com/jltsiren/gbz-base and https://crates.io/crates/gbz-base under the MIT license.

20
How Robust are Multispecies Coalescent Species Delimitations in Taxonomically Complex Systems? A Genomic Assessment Using Mediterranean Tethya Sponges

van der Sprong, J.; Cardone, F.; Hoehna, S.; Schaetzle, S.; Deister, F.; Erpenbeck, D.; Woerheide, G.; Vargas, S.

2026-07-05 evolutionary biology 10.64898/2026.07.04.735074 medRxiv
Top 0.3%
1.1%
Show abstract

Reliable species delimitation underpins biodiversity assessment but remains difficult for organisms with plastic morphology and few diagnostic characters. Multispecies coalescent (MSC) methods can delimit species from genomic data, yet they are rarely tested in taxonomically complex, marine invertebrate groups where they are arguably most needed. We used the three Mediterranean species of the genus Tethya, a rare, well-characterised system within the otherwise taxonomically difficult phylum Porifera-distinguished by multiple independent morphological and ecological characters-to evaluate how robust MSC-based delimitation is in such groups. Analysing 64 single-copy nuclear loci in BEAST2 and BPP, we compared constrained, hypothesis-testing approaches (BFD*, BFdriver, A10) with freer, heuristic ones (SPEEDEMON, A11), and examined their sensitivity to data type, clock model, priors, and the species-collapse threshold. All methods recovered the three recognised Mediterranean species, but the resolution of within-lineage structure was method-dependent. The hypothesis-testing approaches consistently supported six lineages, robustly across data types and model assumptions, whereas the heuristic approaches proved less stable. Configurations without a priori species hypotheses often failed to converge or were computationally intractable, a problem compounded by the relaxed clock. In SPEEDEMON the outcome changed with the collapse threshold. Because our system lacks an independent reference point to calibrate this threshold, any delimitation based on it is poorly constrained. We conclude that constrained, hypothesis-testing delimitation is the most robust and reproducible MSC approach, yielding a quantitative, model-based hypothesis that can be weighed against other lines of evidence to inform taxonomic decisions. By clarifying how these methods behave and how their outcomes should be interpreted, our study offers a practical guide for researchers working on comparably complex systems.